Papers with machine translation training

3 papers
Towards the First NLP Benchmark for Ladin - an Extremely Low-Resource Language (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) are limited in low-resource languages due to lack of labeled training data.
Approach: They propose to use Ladin as a model for sentiment analysis and question answering by incorporating Italian data into machine translation training.
Outcome: The proposed method improves on existing Italian–Ladin translation baselines.
Machine Translation Models are Zero-Shot Detectors of Translation Direction (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to detect the translation direction of parallel text are lacking in the machine translation community.
Approach: They propose an unsupervised approach to detection of translation direction of parallel texts . they use a simple hypothesis that p(translation|original)>p(original|translation) they confirm the approach is effective for high-resource language pairs .
Outcome: The proposed approach achieves document-level accuracies of 82–96% for NMT-produced translations and 60–81% for human translations, based on the model used.
A New Massive Multilingual Dataset for High-Performance Language Technologies (2024.lrec-main)

Copied to clipboard

Challenge: a new massive multilingual dataset is available for language modeling and machine translation training.
Approach: They present a massive multilingual dataset using web crawls from the Internet Archive and CommonCrawl . they use open-source software tools and high-performance computing to acquire, manage and process large corpora .
Outcome: The HPLT language resources is a massive multilingual dataset . it includes monolingual and bilingual corpora extracted from CommonCrawl and the Internet Archive . the results are published online at the journal journal cense4 .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations